Papers by Phillip Benjamin Ströbel
Evaluation of HTR models without Ground Truth Material (2022.lrec-1)
Copied to clipboard
Phillip Benjamin Ströbel, Martin Volk, Simon Clematide, Raphael Schwitter, Tobias Hodel, David Schoch
| Challenge: | Optical Character Recognition (OCR) is a well-established technique for digitising historical printed collections in libraries and archives. |
| Approach: | They propose to use masked language models to evaluate handwritten text recognition models . they propose to introduce GT-free metrics to evaluate models to ensure best results . |
| Outcome: | The proposed model evaluations are based on lexicon-based and masked language models. |
Language Resources for Historical Newspapers: the Impresso Collection (2020.lrec-1)
Copied to clipboard
| Challenge: | digitization efforts are slowly but steadily contributing an increasing amount of facsimiles of cultural heritage documents. |
| Approach: | They propose to use a collection of newspaper data sets composed of text and image resources, curated and published within the context of the ‘impresso - Media Monitoring of the Past’ project. |
| Outcome: | The aim of the impresso resource collection is to contribute to historical language resources, and strengthen approaches to non-standard inputs and foster efficient processing of historical documents. |
How Much Data Do You Need? About the Creation of a Ground Truth for Black Letter and the Effectiveness of Neural OCR (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent advances in Optical Character Recognition and Handwritten Text Recognition have led to more accurate text recognition of historical documents. |
| Approach: | They propose to build a ground truth for a German-language newspaper published in black letter . they also evaluate the performance of different OCR engines and estimate how much data is needed to achieve high-quality OCR results. |
| Outcome: | The proposed model can recognise black letter text and performs well on data they have not seen during training. |